Papers by Nanda Putri Romadhona
BRCC and SentiBahasaRojak: The First Bahasa Rojak Corpus for Pretraining and Sentiment Analysis Dataset (2022.coling-1)
Copied to clipboard
| Challenge: | Code-mixing is prevalent in multilingual societies and is challenging to train . we use data augmentation to build a model to deal with code-mixed inputs . |
| Approach: | They propose to train a model to deal with code-mixing phenomena of Bahasa Rojak using data augmentation to construct a Bahasan Rojakin corpus and a pre-trained model to process input tokens. |
| Outcome: | The proposed model can tag the language of the input token automatically to process code-mixing input. |